Tag
18 articles
This article explains the concept of prompt injection, a technique where hidden instructions are embedded in messages to manipulate AI systems. It uses a court case as an example to show how this can happen and why it matters.
Learn how hidden text in PDFs can trick AI tools like Atlassian's Rovo into stealing sensitive data without user knowledge.
Anthropic's Opus 5, combined with Auto Mode, shows zero success rate in preventing browser-based prompt injection attacks, a major AI security vulnerability.
Security researchers have developed a new technique called 'context bombing' that prevents malicious AI agents from executing harmful actions by overwhelming them with excessive contextual data.
OpenAI's internal AI red-teaming model, GPT-Red, outperformed human red-teamers 84% to 13% on prompt injection tests and discovered novel attack techniques.
OpenAI introduces GPT-Red, an automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. This innovation represents a significant advancement in AI safety research.
Learn how to implement defensive prompt injection techniques using context bombing to protect AI systems from malicious manipulation.
OpenAI introduces ChatGPT's Lockdown Mode to protect sensitive data from prompt injection attacks by disabling web access and research features.
Learn about prompt injection attacks and how OpenAI's new Lockdown Mode aims to protect sensitive data in AI systems.
A simple GitHub issue could have compromised Anthropic’s Claude Code action, exposing projects that use it to potential data breaches and unauthorized access.
A developer has revealed how a malicious code addition in the popular Java library jqwik could have instructed AI coding agents to delete application output, highlighting serious security vulnerabilities in AI-assisted development.
Google warns that malicious web pages are poisoning enterprise AI agents through indirect prompt injections, exploiting hidden HTML code to manipulate AI systems.